The other side of AI safety
13 Jul 2026
Award-winning LMU position paper highlights how AI safety methods can also be used for censorship and manipulation.
13 Jul 2026
Award-winning LMU position paper highlights how AI safety methods can also be used for censorship and manipulation.
Artificial intelligence (AI) is designed not to answer certain questions—for example, those that could facilitate the construction of bombs or other dangerous activities. Technical safety mechanisms are intended to prevent such harmful assistance. However, the same methods can also be misused to suppress legitimate questions about history, politics, or other sensitive topics.
This so-called dual-use risk is the focus of a position paper by LMU researchers Sarah Ball and Dr. Phil Hackemann, which received an award at the International Conference on Machine Learning (ICML) 2026 in Seoul, South Korea, one of the world's leading conferences on machine learning.
“As AI researchers, we must recognize that the methods we develop can also be misused,” says Sarah Ball. “We hope to raise awareness of this risk—not only within the research community, but beyond it as well.”
Ball is a doctoral researcher at LMU's Social Data Science and AI Lab and the Munich Center for Machine Learning. Together with political scientist Phil Hackemann, she works at the intersection of artificial intelligence, AI safety, social data science, and political science.
As AI researchers, we must recognize that the methods we develop can also be misused. We hope to raise awareness of this risk—not only within the research community, but beyond it as well.Sarah Ball
Sarah Ball
“Authoritarian regimes and other powerful actors are already using AI technologies to manipulate facts and, in turn, shape public opinion,” says Phil Hackemann. “That is why it is essential to recognize the dual-use risks inherent in these technologies.”
In their award-winning position paper, The Alignment Community is Unintentionally Building a Censor's Toolkit, Ball and Hackemann challenge the widespread assumption that value alignment—the process of aligning AI systems with predefined values and objectives—necessarily serves the public good. Instead, they argue that the same techniques can also be used to enable censorship and manipulation.
“We need to understand that safety mechanisms designed to prevent misuse can themselves be misused,” Ball explains. "“A model that reliably blocks dangerous content can, using the very same methods, also be prevented from providing legitimate information.”
Hackemann adds: “The tools that make AI safe are also the tools that can turn it into a censor.”
Phil Hackemann | © Benjamin Jenak
According to the researchers, AI systems can be influenced at several stages of their development and deployment.
Even before training begins, specific content can be removed from training datasets. During training, models can be shaped through human feedback, evaluation systems, or predefined guidelines to provide—or refuse—certain types of responses. Even after deployment, system prompts and additional filtering mechanisms can determine which answers an AI system ultimately delivers.
“When an AI system refuses to answer a question, it may not be because it lacks the necessary knowledge,” says Hackemann. “It may instead have been deliberately aligned to withhold certain information.”
As examples, the researchers point to Chinese AI models such as DeepSeek, which do not freely answer politically sensitive questions, including those concerning the 1989 Tiananmen Square crackdown. They also cite regulations issued by China's Cyberspace Administration (CAC), which require AI providers to ensure that their models generate responses deemed “safe” and refuse certain requests.
Another example is the Chinese model Yi-large, which, according to the paper, was observed filtering responses while they were being generated in order to suppress critical content. For Ball and Hackemann, these cases illustrate how techniques originally developed to improve AI safety can also be used to enforce political objectives.
When an AI system refuses to answer a question, it may not be because it lacks the necessary knowledge. It may instead have been deliberately aligned to withhold certain information.Phil Hackemann
The researchers emphasize that the risks are not limited to government censorship. Private companies can likewise influence the political or societal perspectives reflected in AI systems through choices about training data, system prompts, or other forms of intervention. As one example, they point to recent debates surrounding Elon Musk's chatbot Grok.
At the same time, Ball and Hackemann outline ways to reduce the potential for misuse. One proposal is to develop more robust evaluation methods capable of detecting censorship and manipulation in AI systems.
“Users should be able to understand what information an AI system provides, withholds, or refuses to answer,” says Ball.
The researchers also argue that a diverse AI ecosystem is essential.
“When only a small number of companies or governments control key AI-based information systems, this poses a risk to public discourse,” says Hackemann. “Those who determine which values an AI system reflects also influence the information people ultimately receive.”